Papers with Universal Dependencies project

6 papers
How Bad are PoS Tagger in Cross-Corpora Settings? Evaluating Annotation Divergence in the UD Project. (N19-1)

Copied to clipboard

Challenge: Using annotation variation principles, Part-of-Speech tagging performance degrades when applied to test sentences that depart from training data.
Approach: They propose to use the annotation variation principle to identify inconsistencies between annotations . they also evaluate their impact on prediction performance .
Outcome: The proposed method can detect errors in gold standard annotations and improve prediction performance.
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)

Copied to clipboard

Challenge: a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part.
Approach: They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers.
Outcome: The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets .
Automatic Extraction of Rules Governing Morphological Agreement (2020.emnlp-main)

Copied to clipboard

Challenge: Creating a descriptive grammar is an indispensable step for language documentation but it is tedious and time-consuming.
Approach: They propose a framework for extracting a first-pass grammatical specification from raw text in a concise, human- and machine-readable format.
Outcome: The proposed framework extracts a grammatical specification that is nearly equivalent to those created with large amounts of gold-standard annotated data.
Universal Dependencies for Western Sierra Puebla Nahuatl (2022.lrec-1)

Copied to clipboard

Challenge: Annotated corpus of western Sierra Puebla Nahuatl conforms to universal dependency project annotation guidelines . morphological and syntactic phenomena can be analyzed quantitatively with a large enough corpus .
Approach: They present a morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl . it is the first indigenous language of Mexico to be added to the Universal Dependencies project . UD is a widely-used annotation framework whose aim is to provide a consistent schema for morphological and syntactic phenomena for all of the world's languages.
Outcome: The morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl conforms to the universal dependency project annotation guidelines.
Data Augmentation via Dependency Tree Morphing for Low-Resource Languages (D18-1)

Copied to clipboard

Challenge: Lack of sizable training datasets leads to poor performance in low-resource languages.
Approach: They propose two techniques to augment training sets of low-resource languages using dependency trees.
Outcome: The proposed methods improve on the training datasets for low-resource languages.
Errator: a Tool to Help Detect Annotation Errors in the Universal Dependencies Project (L18-1)

Copied to clipboard

Challenge: UD project aims to develop cross-linguistically consistent treebank annotations for a wide array of languages.
Approach: They introduce tools that implement the annotation variation principle to help annotators find and correct errors in UD treebanks.
Outcome: The proposed tools can be used to correct errors in UD treebank annotations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations